Papers by Jey Han Lau
Copied to clipboard
| Challenge: | Existing evaluation methods focus on single-round inference, but this view is problematic in real-world applications. |
| Approach: | They propose a framework that couples Steering Token Calibration with Semantic Alignment to ensure that LLMs are correctly aligned across gender, race, and sentiment. |
| Outcome: | The proposed framework outperforms baseline methods in achieving precise distributional control in attribute generation tasks. |
Copied to clipboard
| Challenge: | Existing approaches to interpret task-oriented dialogue systems employ an implicit reasoning strategy that makes the model predictions uninterpretable to humans. |
| Approach: | They propose a neuro-symbolic approach that performs explicit reasoning that justifies model decisions by reasoning chains. |
| Outcome: | The proposed approach achieves better results and introduces an interpretable decision process. |
Copied to clipboard
| Challenge: | a paper examining the influence of document context on acceptability judgements for English sentences is published in journal journal of linguistics. |
| Approach: | They propose to use document context to assess acceptability judgements for English sentences . they also test the accuracy of neural models that incorporate document context during training . |
| Outcome: | The proposed model improves acceptability ratings for ill-formed sentences, but reduces them for well-formed ones. |
Copied to clipboard
| Challenge: | Existing work has addressed each element individually, but this study focuses on LipKey, the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles. |
| Approach: | They propose a novel news dataset that consists of highly absent keyphrases . they combine lips keyphrase and TF-IDF to obtain abstractive summaries . |
| Outcome: | The proposed dataset is the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles. |
Copied to clipboard
| Challenge: | X, Meta, and TikTok are experimenting with community-based factchecking . community-driven verification is a way to provide explanatory notes that clarify why a post might be misleading . |
| Approach: | They propose a framework that optimizes the helpfulness of explanatory notes and the reason for this by automatically optimizing the prompt definitions. |
| Outcome: | The proposed framework improves helpfulness and reason prediction on 104k posts with user-provided notes and helpfulness labels. |
Copied to clipboard
| Challenge: | Despite having the fourth largest speaker population in the world, 1 Indonesian is under-represented in NLP. |
| Approach: | They propose to use a large-scale Indonesian summarization dataset to test extractive and abstractive summarizing methods. |
| Outcome: | The proposed methods are compared with multilingual and monolingual BERT-based models. |
Copied to clipboard
| Challenge: | Existing summarisation systems are not up to such complex tasks, yet limited tools exist to determine where and why they are failing. |
| Approach: | They propose to use a dataset to evaluate the quality of summarisation systems in the biomedical domain. |
| Outcome: | The proposed model can be used to evaluate the quality of summarisation systems in the biomedical domain. |
Copied to clipboard
| Challenge: | Social media rumours can cause significant economic and social disruption. |
| Approach: | They propose a rumour detection algorithm that leverages transformers and graph attention networks to jointly model social media conversations and the network of users who engaged in them. |
| Outcome: | The proposed algorithm produces superior performance over four widely used benchmark rumour datasets in English and Chinese. |
Copied to clipboard
| Challenge: | In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks. |
| Approach: | They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia. |
| Outcome: | The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons. |
Copied to clipboard
| Challenge: | Recent advances in abstractive text summarization have created plausible summaries, but it is unclear if they truly possess the capability of information consolidation to generate summary. |
| Approach: | They propose to prompt large language models to generate meta-reviews and use evaluation metrics to assess the quality of generated meta- reviews. |
| Outcome: | The proposed framework proves that human meta-reviewers follow a framework of sentiment consolidation to write meta- reviews compared with prompting them with simple instructions. |
Copied to clipboard
| Challenge: | Existing models generate fluent and coherent summaries, but inconsistencies can be found in generated summary. |
| Approach: | They propose to use symbolic knowledge distillation to improve the factual consistency of smaller pretrained models for dialogue summarization. |
| Outcome: | The proposed model outperforms baseline models in BART, PEGASUS, and Flan-T5 in factual consistency and accuracy. |
Copied to clipboard
| Challenge: | There are more than 700 languages spoken in Indonesia, equal to 10% of the world's languages, second only to Papua New Guinea. |
| Approach: | They focus on the languages spoken in Indonesia, the world's second most linguistically diverse nation, and the fourth most populous nation of the world. |
| Outcome: | The proposed model is based on the languages spoken in Indonesia, the world's second-most linguistically diverse nation, with 273 million people spread over 17,508 islands. |
Copied to clipboard
| Challenge: | Existing studies on the impact of human label variation on model fairness have not explored the interaction between HLV and performance. |
| Approach: | They compare human label variation (HLV) training methods with four other methods . they find that HLV methods improve performance without harming fairness . |
| Outcome: | The proposed methods improve fairness without explicit debiasing under certain configurations. |
Copied to clipboard
| Challenge: | Existing evaluation approaches to multi-document summarization of biomedical literature lack consistency and transparency. |
| Approach: | They propose a systematic approach to human evaluation of biomedical summaries and apply it to analyze the summary generated by two current evaluation models. |
| Outcome: | The proposed evaluation framework is based on two state-of-the-art models and examines the summaries generated by the two models to understand the deficiencies of existing evaluation approaches. |
Copied to clipboard
| Challenge: | Existing methods for paraphrasing multiword expressions in context are unsupervised . multiwords are notoriously difficult to model because the meaning of the whole can diverge substantially from that of the component words. |
| Approach: | They propose an unsupervised approach to paraphrasing multiword expressions in context using monolingual corpus data and pre-trained language models. |
| Outcome: | The proposed method outperforms all unsupervised systems and rivals supervised systems on the SemEval 2022 idiomatic text similarity task. |
Copied to clipboard
| Challenge: | Negation is an important linguistic phenomenon which denotes non-existence, denial, or contradiction. |
| Approach: | They propose a natural language inference test suite to test models for negation . they use a linguistic framework to analyze negation types and constructions . |
| Outcome: | The proposed test suite is more challenging than existing benchmarks on negation . it includes annotation of negation types and constructions grounded in linguistic theory . |
Copied to clipboard
| Challenge: | Existing EaaS watermarks can be removed by paraphrasing when attackers clone the model. |
| Approach: | They propose a method that integrates a target embedding into the original embeddable based on the presence of trigger words in the input text. |
| Outcome: | The proposed technique is empirically and theoretically robust against paraphrasing. |
Copied to clipboard
| Challenge: | Using a semantic memory, we score each utterance along three interpretable dimensions: Novelty, Relevance, and Implication Scope. |
| Approach: | They propose a framework for Conversational Information Gain that evaluates each utterance in terms of how it advances collective understanding of the target topic. |
| Outcome: | The proposed framework evaluates each utterance in terms of how it advances collective understanding of the target topic. |
Copied to clipboard
| Challenge: | Recent advances in deep neural networks have created applications for a range of different domains. |
| Approach: | They propose a grey-box adversarial attack and defence framework for sentiment classification . they show that the framework produces an improved classifier that is robust in defending . |
| Outcome: | The proposed framework produces an improved classifier that is robust in defending against multiple adversarial attacking methods. |
Copied to clipboard
| Challenge: | Existing studies on rumour detection are concerned with timing, but few are interested in how early we can detect them. |
| Approach: | They propose a method that integrates reinforcement learning to learn the minimum number of posts required before classifying an event as a rumour. |
| Outcome: | The proposed model detects rumours earlier than state-of-the-art systems while maintaining comparable accuracy. |
Copied to clipboard
| Challenge: | Existing AVR benchmarks focus on single-step reasoning, emphasizing the end result but neglecting the multi-stage nature of reasoning process. |
| Approach: | They propose a multi-stage AVR benchmark based on RAVEN to assess reasoning across varying levels of complexity. |
| Outcome: | The proposed metric considers the correctness of intermediate steps in addition to the final outcomes. |
Copied to clipboard
| Challenge: | Existing methods for lexical substitution using pre-trained language models have some limitations. |
| Approach: | They propose an unsupervised method for lexical substitution using pre-trained language models. |
| Outcome: | The proposed method outperforms baseline models and establishes a state-of-the-art without supervision or fine-tuning. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can simulate non-native-like English use observed in human second language (L2) learners interfered with by their native first language (N1) knowledge. |
| Approach: | They use large language models to simulate non-native-like English use observed in human second language (L2) learners, and then compare their outputs to real L2 learner data. |
| Outcome: | The proposed models replicate L1-dependent patterns observed in human second language (L2) learners, with distinct influences from various languages. |
Copied to clipboard
| Challenge: | Discourse analysis is a systematic way to understand how texts are segmented hierarchically into discourse units. |
| Approach: | They propose a top-down approach to discourse parsing that is conceptually simpler than its predecessors. |
| Outcome: | The proposed model eliminates the decoder and reduces the search space for splitting points. |
Copied to clipboard
| Challenge: | Knowledge base question answering (KBQA) aims to answer user questions in natural language using rich human knowledge stored in large KBs. |
| Approach: | They propose a model that injects schema contexts into entity retrieval and logical form generation to enhance generalizability. |
| Outcome: | The proposed model outperforms state-of-the-art models on two commonly used benchmark datasets across a variety of test settings. |
Copied to clipboard
| Challenge: | FLUKE introduces controlled variations across linguistic levels and leverages large language models with human validation to generate modifications. |
| Approach: | They propose a framework for assessing model robustness through systematic minimal variations of test data. |
| Outcome: | The proposed framework evaluates models and LLMs across six diverse NLP tasks and shows that they are more robust to natural, fluent modifications than base models. |
Copied to clipboard
| Challenge: | neutralisation is used to justify lack of action or promote an alternative view of climate change . action on climate change has become an increasingly partisan issue with strong opposition voices discrediting scientists and spreading scepticism and misinformation. |
| Approach: | They propose to use neutralisation techniques to introduce the problem to the nlp community and to collect manual annotations of neutralised techniques in text relating to climate change. |
| Outcome: | The proposed models are supervised and semi-supervised by a team of researchers from the nlp and the ccsc. |
Copied to clipboard
| Challenge: | despite being spoken by 200 million people, the Indonesian language is underrepresented in NLP research. |
| Approach: | They propose a dataset for Indonesian that includes seven NLP tasks . they also propose 'indonesian language evaluation Montage' tasks that are based on previous work . |
| Outcome: | The proposed dataset shows that IndoBERT outperforms IndoLEM over most of the tasks. |
Copied to clipboard
| Challenge: | Existing data on ESL speakers' communication and interaction skills are lacking in the evaluation of the sophisticated features of dialogue. |
| Approach: | They propose an evaluation framework for interactive dialogue assessment in ESL speakers. |
| Outcome: | The proposed framework provides a means to assess ESL communication, useful for language assessment. |
Copied to clipboard
| Challenge: | Using this framework, we annotated 5,657 sentences with human judges and 15,494 sentences with GPT-4o from two domains: TV debates and radio panel discussions. |
| Approach: | They propose an evaluation framework for analyzing the facilitation strategies of moderators across different domains/scenarios by examining their motives (Why), dialogue acts (How) and target speaker (Who). |
| Outcome: | The framework is generalisable across domains and reveals distinct modes of moderation: debate moderators emphasise coordination and facilitate interaction through questions and instructions, panel discussion moderator prioritize information provision and actively participate in discussions. |
Copied to clipboard
| Challenge: | a recent surge of interest in deep learning has led to creative applications for poetry generation . a novel joint architecture captures language, rhyme and meter for sonnet modelling . |
| Approach: | They propose a joint architecture that captures language, rhyme and meter for sonnet modelling. |
| Outcome: | The proposed architecture captures language, rhyme and meter for sonnet modelling. |
Copied to clipboard
| Challenge: | Existing methods for summarizing opinions from large-scale online reviews are not available for crowdsourcing and are difficult to crowdsource. |
| Approach: | They propose a domain-agnostic modular approach guided by review aspects to separate tasks of aspect identification, opinion consolidation, and meta-review synthesis to enable greater transparency and ease of inspection. |
| Outcome: | The proposed approach generates more grounded summaries than baseline models, as verified through automated and human evaluations. |
Copied to clipboard
| Challenge: | Existing work on probing of pretrained language models has focused on sentence-level syntactic tasks. |
| Approach: | They introduce document-level discourse probing to evaluate the ability of pretrained LMs to capture document- level relations. |
| Outcome: | The proposed model performs best in encoder, but only in the encoder layer. |
Copied to clipboard
| Challenge: | a recent study shows that context affects our perception of sentence acceptability, but few studies investigate how it affects language models. |
| Approach: | They compare acceptability ratings of sentences judged in isolation with a relevant context and with an irrelevant context. |
| Outcome: | The proposed model achieves state-of-the-art for unsupervised acceptability prediction. |
Copied to clipboard
| Challenge: | Existing studies on second language (SL) assessment of conversational fluency and interactivity have focused on written correction or pronunciation from ASR. |
| Approach: | They propose a framework that assesses the relationships between micro-level linguistic features and macro-level interactivity labels for Chinese-as-a-second-language dialogues. |
| Outcome: | The proposed framework is interpretable and can be adapted to other languages for second-language dialogue evaluation. |
Copied to clipboard
| Challenge: | In IndoBERTweet, a pretraining model for Indonesian Twitter is extended with domain-specific vocabulary. |
| Approach: | They propose a pretraining model that extends a monolingual Indonesian BERT model with domain-specific vocabulary. |
| Outcome: | The proposed model can be initialized with the average BERT subword embedding five times faster than existing methods for vocabulary adaptation. |
Copied to clipboard
| Challenge: | Task-oriented dialogue models can learn non-transferable generalizations by using shortcuts in the data. |
| Approach: | They propose a contrastive learning framework to encourage models to ignore cues and focus on generalisable patterns. |
| Outcome: | The proposed framework performs exceptionally well on task-oriented dialogue datasets. |
Copied to clipboard
| Challenge: | Existing tools for ESL assessment focus on writing skills and lack in support for dynamic spoken interactions. |
| Approach: | They propose an approach that integrates automatic ESL dialogue assessment and a framework that categorizes moderation strategies to assess conversational engagement and moderation effectiveness. |
| Outcome: | The proposed approach integrates automatic ESL dialogue assessment and categorizes moderation strategies. |
Copied to clipboard
| Challenge: | Existing fact-checking systems struggle with attribution quality, as their generated explanations can include hallucinations. |
| Approach: | They propose a protocol to assess attribution quality in fact-checking explanations using human annotation and automatic annotation. |
| Outcome: | The proposed protocol can be automated, the authors show . best-performing LLMs still generate explanations that are not always accurate . |
Copied to clipboard
| Challenge: | Existing work on factual inconsistency in abstractive summarization addresses this problem. |
| Approach: | They propose a dataset with fine-grained factual error annotations named DIASUMFACT and an unsupervised model named ENDERANKER. |
| Outcome: | The proposed model performs on par with the state-of-the-art models while requiring fewer resources. |
Copied to clipboard
| Challenge: | Using multilingual summarization evaluation methods is more reliable and interpretable than manual methods. |
| Approach: | They propose to use multilingual BERT within BERTScore to evaluate summarization evaluation metrics . they use English datasets that are not representative of modern summarizing systems . |
| Outcome: | The proposed methods perform well across all languages, at a level above that for English. |
Copied to clipboard
| Challenge: | Recent VSE models combine simple pooling methods with hard triplet loss to improve performance. |
| Approach: | They propose an adaptive pooling strategy that allows the model to learn how to aggregate features through a combination of simple pooling methods. |
| Outcome: | The proposed strategy outperforms current state-of-the-art systems on image-to-text and text-toimage retrieval. |
Copied to clipboard
| Challenge: | Existing studies on explainable fake news or rumour detection by and large use attention weights as explanation, but the use of attention weighted explanations is problematic. |
| Approach: | They propose a causal mediation analysis approach to explain the decision-making process of neural models for rumour detection on Twitter by identifying salient tweets that explain model predictions and highlighting causally impactful words in the tweets. |
| Outcome: | The proposed approach shows strong agreement with human judgements for critical tweets determining the truthfulness of stories. |
Copied to clipboard
| Challenge: | Topic coherence is increasingly being used to evaluate topic models and filter topics for end-user applications. |
| Approach: | They propose to use topic intrusion to guess an outlier topic given a document and a few topics to automate the task. |
| Outcome: | The proposed method improves upon the state-of-the-art method and shows it can be used as an alternative to topic perplexity evaluation. |
Copied to clipboard
| Challenge: | a paper on automatic sentencing was a source of debate at EMNLP 2019 . paper examines whether particular datasets and tasks should be off-limits for NLP research . |
| Approach: | They propose a neural model which performs structured prediction of individual charges laid against an individual and the prison term associated with each. |
| Outcome: | The proposed model can predict the prison term associated with a given case on a large-scale dataset of real-world Chinese court cases. |